Papers with evaluation toolkit
ARQA: A Benchmark for Grounded Table–Text QA in Enterprise Annual Reports (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing QA benchmarks focus on retrieval or single-modality reasoning . annual reports are a company's definitive record of performance . |
| Approach: | They propose an annual report QA benchmark that compares QAs with lookups, arithmetics, and insights. |
| Outcome: | The proposed benchmarks show strong factual retrieval but persistent weaknesses in grounded arithmetic and causal reasoning. |
Analyzing and Evaluating Faithfulness in Dialogue Summarization (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on faithfulness of text summarization have not been conducted on abstractive summarizing. |
| Approach: | They propose a method to evaluate faithfulness of dialogue summarization models by multi-choice questions. |
| Outcome: | The proposed method can facilitate the development of dialogue summarization systems. |
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Queries (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) generate code for productive activities, but current benchmarks for code synthesis are oriented towards introductory tasks on algorithm and data science. |
| Approach: | They propose a code benchmark to mirror the complexity and variety of scenarios in real-world coding tasks. |
| Outcome: | The proposed benchmark improves on 39 large language models with close HumanEval scores and achieves an efficiency increase of more than 4 times. |
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive results across a broad array of tasks, yet their capacity for complex, domain-specific mathematical reasoning remains underexplored. |
| Approach: | They propose a benchmark to evaluate Large Language Models on mathematical modeling challenges to wireless communications engineering. |
| Outcome: | The proposed benchmark evaluates LLMs on mathematical modeling challenges to wireless communications engineering. |